.gguf files from HuggingFace and pass them to --model.
DeepSeek-V3 and DeepSeek-R1 use Multi-head Latent Attention (MLA). Pass
-mla 3
(the default) for best performance. Lower values reduce memory use at a speed cost.